Papers with Topic modeling

14 papers
Topic Modeling in Embedding Spaces (2020.tacl-1)

Copied to clipboard

Challenge: Existing topic models fail to learn interpretable topics when working with large and heavy-tailed vocabularies.
Approach: They propose an embedded topic model that integrates word embeddings with a categorical distribution that is the natural parameter between the word’s embeddment and an embeddement of its assigned topic.
Outcome: The embedded topic model outperforms existing topic models in terms of topic quality and predictive performance.
STREAM: Simplified Topic Retrieval, Exploration, and Analysis Module (2024.acl-short)

Copied to clipboard

Challenge: Topic modeling is a widely used technique to analyze large document corpora.
Approach: They propose a module for topic retrieval, exploration, and analysis that implements multiple intruder-word based topic evaluation metrics.
Outcome: The proposed module implements multiple intruder-word based topic evaluation metrics and extends existing datasets.
Neural Topic Modeling with Large Language Models in the Loop (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated promising capabilities in topic discovery, but their direct application to topic modeling suffers from issues such as incomplete topic coverage, misalignment of topics, and inefficiency.
Approach: They propose a novel LLM-in-the-loop framework that integrates Large Language Models with Neural Topic Models (NTMs) global topics and document representations are learned through the NTM, while an LLM refines these topics using an Optimal Transport (OT)-based alignment objective.
Outcome: The proposed framework improves topic interpretability while preserving the efficiency of existing NTMs.
A Query-Driven Topic Model (2021.findings-acl)

Copied to clipboard

Challenge: Topic modeling is an unsupervised method for revealing the hidden semantic structure of a corpus.
Approach: They propose a query-driven topic model that allows users to specify a simple query in words or phrases and return query-related topics.
Outcome: The proposed model is particularly attractive when the query has a low occurrence in a text corpus, making it difficult for traditional topic models to identify relevant topics.
TopicGPT: A Prompt-based Topic Modeling Framework (2024.naacl-long)

Copied to clipboard

Challenge: TopicGPT uses large language models to uncover latent topics in text . topic models represent topics as bags of words that require "reading the tea leaves" topic models also offer limited control over formatting and specificity of topics .
Approach: TopicGPT uses large language models to uncover latent topics in text . authors propose a prompt-based framework that produces topics that align better with human categorizations .
Outcome: TopicGPT produces topics that align better with human categorizations compared to competing methods.
Tree-Structured Topic Modeling with Nonparametric Neural Variational Inference (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for topic modeling learn topics with a flat structure . however, such methods have data scalability issues .
Approach: They propose to use nonparametric neural variational inference to extract a tree-structured topic model with reasonable structure, low redundancy, and adaptable widths.
Outcome: The proposed model extracts a tree-structured topic hierarchy with reasonable structure, low redundancy, and adaptable widths.
MirasText: An Automatically Generated Text Corpus for Persian (L18-1)

Copied to clipboard

Challenge: Natural language processing is one of the most important fields of artificial intelligence.
Approach: They propose to use MirasText to generate Persian text corpus from Persian websites . MiraSText has over 2.8 million documents and over 1.4 billion tokens .
Outcome: The generated corpus has over 2.8 million documents and over 1.4 billion tokens . MirasText has over 800 billion token tokens and more than 300 thousand articles .
Understanding Cross-Domain Adaptation in Low-Resource Topic Modeling (2025.acl-long)

Copied to clipboard

Challenge: Existing topic modeling models struggle in low-resource settings where data is limited . et al., 2003: domain adaptation for low-source topic modeling is challenging in low resources .
Approach: They propose a domain adaptation framework that disentangles domaininvariant and domain-specific components to improve topic adaptation.
Outcome: The proposed model outperforms state-of-the-art methods on low-resource datasets on diverse datasets.
GINopic: Topic Modeling with Graph Isomorphism Network (2024.naacl-long)

Copied to clipboard

Challenge: Recent studies focus on the representation of documents as a sequence of words, but word dependency patterns are not captured in topic modeling.
Approach: They propose a topic modeling framework based on graph isomorphism networks to capture word dependencies between words.
Outcome: The proposed framework is compared with existing topic models on a dataset of a large text collection and shows that it can uncover the underlying topics in an unsupervised manner.
Neural Topic Modeling via Contextual and Graph Information Fusion (2025.emnlp-main)

Copied to clipboard

Challenge: Existing topic models generate uninformative and incoherent topics that hinder interpretable insights from managing textual data.
Approach: They propose to incorporate contextual and graph information to improve the variational autoencoder framework by combining contextual and bag-of-words information.
Outcome: The proposed framework generates more coherent and diverse topics on three benchmark datasets and achieves strong performance on automatic and manual evaluations.
LLM-Guided Semantic-Aware Clustering for Topic Modeling (2025.acl-long)

Copied to clipboard

Challenge: Experimental results show that topic modeling is competitive compared to closed-source methods.
Approach: They propose a semi-supervised topic modeling method that combines LLMs with clustering to improve topic generation and distribution.
Outcome: The proposed method outperforms state-of-the-art methods that utilize GPT-4 on topic alignment and exhibits competitive performance compared to Neural Topic Models on topic quality.
Enhancing Short-Text Topic Modeling with LLM-Driven Context Expansion and Prefix-Tuned VAEs (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing topic models often lack sufficient word co-occurrence in short texts, resulting in incoherent topics.
Approach: They propose to use large language models to extend short texts into more detailed sequences before applying topic modeling to solve semantic inconsistency problem.
Outcome: The proposed approach significantly outperforms current state-of-the-art topic models on real-world datasets with extreme data sparsity.
Semantic Component Analysis: Introducing Multi-Topic Distributions to Clustering-Based Topic Modeling (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for topic modeling fail to scale to large datasets or assume one topic per document.
Approach: They propose a topic modeling technique that discovers multiple topics per sample . they evaluate SCA on Twitter datasets in English, Hausa and Chinese .
Outcome: The proposed technique outperforms the LLM-based TopicGPT on Twitter datasets with similar compute budgets.
CobwebTM: Probabilistic Concept Formation for Lifelong and Hierarchical Topic Modeling (2026.findings-acl)

Copied to clipboard

Challenge: Topic modeling seeks to uncover latent semantic structure in text corpora with minimal supervision.
Approach: They propose a lifelong hierarchical topic model based on incremental probabilistic concept formation that constructs semantic hierarchies online without predefining the number of topics.
Outcome: The proposed model achieves strong topic coherence, stable topics over time, and high-quality hierarchies without predefining the number of topics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations